Papers by Sai Sathiesh Rajan

1 papers
Localizing Malicious Outputs from CodeLLM (2025.findings-emnlp)

Copied to clipboard

Challenge: Using FreqRank, we localize malicious components in outputs for triggered inputs and their corresponding backdoor triggers.
Approach: They propose a mutation-based defense to localize malicious components in LLM outputs and their corresponding backdoor triggers.
Outcome: The proposed defense has an average attack success rate (ASR) of 86.6% and can localize the backdoor triggers in 98% of cases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations